Repository navigation
feat: allow EKS Auto Mode nodes to reach ElastiCache/RDS/RDS-proxy - #63
Merged
Merged
Conversation
EKS Auto Mode nodes attach the cluster PRIMARY security group (module.eks.cluster_primary_security_group_id), while managed node groups use the module's node shared SG (node_security_group_id). The data-layer SG ingress rules only referenced the managed node SG, so any pod on an Auto Mode node was blocked from Redis (6379) and MySQL (3306). The module already handles this SG split for node<->node coexistence (comet_eks/main.tf) but never extended it to the data layer. Symptom: on a cluster with an Auto Mode nodepool (e.g. stsaasuat's arm64 `multiarch` pool), a pod that lands there can't reach ElastiCache — e.g. the mysql-db-migration sync hook's waitForResources loops "redis timeout" forever and wedges the ArgoCD sync. - comet_eks: expose `cluster_primary_security_group_id` output (the SG Auto Mode nodes use). - comet_elasticache / comet_rds: add `*_auto_mode_allow_from_sg` var (string, default null) + a second ingress rule (Redis 6379 / MySQL 3306) created only when it's set. Mirrors the module's existing coexistence-rule pattern. - rds_proxy: allowed_sg_ids already list/for_each — root now passes both the managed node SG and (when Auto Mode is enabled) the cluster primary SG. - root main.tf: wire all three, guarded by `var.enable_eks && var.eks_enable_auto_mode`. Default null / Auto-Mode-off => ZERO diff for existing clusters. Suggested release: v5.5.0. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
This was referenced Aug 13, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
User description
Problem
EKS Auto Mode nodes attach the cluster primary security group (
module.eks.cluster_primary_security_group_id, e.g.sg-04eea458…/eks-cluster-sg-<cluster>), while managed node groups use the module's node shared SG (node_security_group_id, e.g.sg-00c20cef…). The data-layer SG ingress rules (ElastiCache 6379, RDS 3306) referenced only the managed node SG — so any pod that lands on an Auto Mode node is blocked from Redis/MySQL.The module already knows these SGs are distinct (
comet_eks/main.tfadds node↔cluster coexistence rules gated onenable_auto_mode) but never extended that to the data layer.Observed on stsaasuat: after adding the arm64
multiarchAuto Mode nodepool, themysql-db-migrationsync hook landed there, itswaitForResourcesloopedredis timeoutforever → never completed → the whole comet-ml ArgoCD app wedged "waiting for completion of hook … mysql-db-migration". (S3 is fine — Gateway VPC endpoint, route-based.)Change
cluster_primary_security_group_id(the SG Auto Mode nodes attach; previously unexposed).*_auto_mode_allow_from_sgvariable (string, defaultnull) + a second ingress rule (Redis 6379 / MySQL 3306) created only when the var is set. Mirrors the existing single-SG rule.allowed_sg_idsis already list/for_each — root now passes both the managed node SG and (when Auto Mode is on) the cluster primary SG.var.enable_eks && var.eks_enable_auto_mode.Safety
Default
null/ Auto-Mode-off → zero diff for every existing non-Auto-Mode cluster. On an Auto Mode cluster,terraform planshows exactly: +1 ElastiCache ingress rule, +1 RDS ingress rule, and the RDS-proxy allow-list gaining the cluster SG — 0 destroy.Suggested release: v5.5.0. stsaasuat bumps
?refto it + applies → Auto Mode pods reach Redis/MySQL, the wedged migration hook completes, and the sync (incl. the dply-utils 2.8.0 fix) unblocks.🤖 Generated with Claude Code
Generated description
Below is a concise technical summary of the changes proposed in this PR:
Enable EKS Auto Mode nodes to access ElastiCache, RDS, and the RDS proxy by exposing the cluster primary security group and wiring it into data-layer access rules. Add conditional Redis/MySQL ingress rules and extend the proxy allow-list while preserving zero changes when Auto Mode is disabled.
cluster_primary_security_group_idfrom the EKS module so the root configuration can conditionally authorize Auto Mode nodes.Modified files (1)
Latest Contributors(2)
Modified files (5)
Latest Contributors(2)